Papers by Dota Tianai Dong
Using Perspectival Words Is Harder Than Vocabulary Words for Humans —and Even More So for Multimodal Language Models (2026.acl-long)
Copied to clipboard
| Challenge: | Existing evaluations of multimodal language models focus on vocabulary words with relatively stable, context-independent meanings in conversation, such as object names, colors, and verbs. |
| Approach: | They compare human and multimodal language models in their use of three word types: vocabulary, possessives, and demonstratives. |
| Outcome: | The models approach human-level performance on using vocabulary, but exhibit clear deficits with possessives and even greater difficulties with demonstratives. |